Papers with public benchmarks LRS3
MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech Recognition (2023.acl-long)
Copied to clipboard
| Challenge: | Audio-visual speech recognition (AVSR) leverages multimodal signals to understand human speech. |
| Approach: | They propose an adversarial network to refine frame-level modality-invariant representations to bridge the distribution gap between modalities. |
| Outcome: | The proposed approach outperforms the state-of-the-art on public benchmarks LRS3 and LRS2 on the modalities of AVSR. |
Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition (2023.acl-long)
Copied to clipboard
| Challenge: | Existing efforts to improve robustness of audio-visual speech recognition with visual information focus on audio modality . current approaches introduce noise adaptation techniques to improve reliability of AVSR task . |
| Approach: | They propose a visual-invariant modality to strengthen robustness of audio-visual speech recognition (AVSR) it can adapt to any testing noises without dependence on noisy training data, a.k.a., unsupervised noise adaptation. |
| Outcome: | The proposed method outperforms existing state-of-the-arts on visual speech recognition task under various noisy and clean conditions. |